Papers by Sanjay Krishna Gouda

3 papers
BASS: Batched Attention-optimized Speculative Sampling (2024.findings-acl)

Copied to clipboard

Challenge: Speculative decoding has emerged as a powerful method to improve latency and throughput in hosting large language models.
Approach: They propose a batched speculative decoding system that generates sequences at an average speed of 5.8ms per token and a batch size of 8 at a 2.15 speed-up over optimized regular decoding.
Outcome: The proposed system achieves state-of-the-art latency and speed-up over optimized regular decoding.
Token Alignment via Character Matching for Subword Completion (2024.findings-acl)

Copied to clipboard

Challenge: Generative models struggle with prompts corresponding to partial tokens due to tokenization, where partial token is out-of-distribution during inference.
Approach: They propose a method to alleviate tokenization artifact on text completion by backtracking to the last complete tokens and aligning subsequent generations to match with the prompt.
Outcome: The proposed method shows that it improves on partial token scenarios with only a minor time increase.
CodeFort: Robust Training for Code Generation Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing research efforts to improve code generation models are inadequate . code generation model performance is degraded under small perturbations .
Approach: They propose a framework to improve the robustness of code generation models by generalizing code perturbations to enrich training data and enabling various robust training strategies.
Outcome: The proposed framework increases pass rates and robustness drop rate against code-syntax perturbations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations